Skip to content

feat: add maintainer-triggered PR description assessment - #4902

Merged
KSchlobohm merged 5 commits into
github:mainfrom
KSchlobohm:kschlobohm-oct-9-upstream-assessment
Oct 10, 2026
Merged

KSchlobohm merged 5 commits into
github:mainfrom
KSchlobohm:kschlobohm-oct-9-upstream-assessment

Conversation

@KSchlobohm

@KSchlobohm KSchlobohm commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

Description

Adds an optional workflow that collaborators with write access or higher trigger with the pr-assess label. It compares the PR description with the code changes, flags material omissions or contradictions, and posts a concise assessment with an outcome label to help reviewers spot gaps.

Testing

I tested outcome-label handling on my fork. All four scenarios passed: first assessment, changed verdict, unchanged verdict with a new comment, and cleanup of two stale outcome labels. Suggested corrections were limited to the PR description.

Outcome labels follow the existing extension-submission remove/add pattern; a matching outcome is left unchanged.

Title stability was also tested on the fork: a stable-title control returned the normal verdict, and a title-change retry returned inconclusive and explained the change. The title now joins head, base, and body in the final input check. An earlier attempt did not change the title and was not evidence of the safeguard.

Automated check: .\.venv\Scripts\python.exe -m pytest tests\test_github_workflows.py -q — 136 passed, 16 skipped. gh aw compile pr-assess --strict --validate --no-check-update passed with the expected pull_request_target warning. uv tool run ruff check tests\test_github_workflows.py, npm exec --offline -- markdownlint-cli2 docs\guides\agentic-sdlc.md, and git diff --check passed.

  • Tested locally with uv run specify --help
  • Ran existing tests with uv sync && uv run pytest
  • Tested with a sample project (if applicable)

AI Disclosure

  • I did not use AI assistance for this contribution
  • I did use AI assistance (fill in the disclosure below)

AI disclosure: Prepared and created with GitHub Copilot (GPT-6.1 Sol) under human supervision.

Port the complete pr-assess workflow with concise reviewer-facing comments,
bounded outcome-label updates, focused tests, and usage guidance.

Keep the reviewed gh-aw v0.89.21 runtime pin isolated from existing workflows.

Assisted-by: GitHub Copilot (model: GPT-6.1 Sol, autonomous)
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 9ee0ab16-f074-4303-82b9-d11bfad16175
Copilot AI balanced review requested due to automatic review settings October 9, 2026 19:28

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Authorization exceeds the stated maintainer scope, required labels are not provisioned, and the AI disclosure is incomplete.

3 open findings
What changed in this PR

Adds a maintainer-triggered workflow that evaluates PR description alignment and reports bounded comments and outcome labels.

Changes:

  • Adds the pr-assess agentic workflow and generated lock file.
  • Adds focused workflow configuration and prompt-contract tests.
  • Documents usage and pins the isolated gh-aw runtime dependency.
File Description
.github/​workflows/​pr-assess.md Defines assessment behavior and safeguards.
.github/​workflows/​pr-assess.lock.yml Provides the compiled GitHub Actions workflow.
.github/​aw/​actions-lock.json Pins gh-aw setup v0.89.21.
tests/​test_github_workflows.py Tests triggers, permissions, outputs, and reporting.
docs/​guides/​agentic-sdlc.md Documents maintainer usage and outcomes.

🧠 Review effort: Balanced


💡 Configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread .github/workflows/pr-assess.md
Comment thread .github/workflows/pr-assess.md
Comment thread .github/workflows/pr-assess.md
@KSchlobohm
KSchlobohm requested a balanced review from Copilot October 9, 2026 19:46
@KSchlobohm
KSchlobohm marked this pull request as ready for review October 9, 2026 19:47
@KSchlobohm
KSchlobohm requested a review from mnriem as a code owner October 9, 2026 19:48

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Separate label removal and addition can leave incorrect verdict state after partial failures.

3 open findings
1 resolved since last review

🧠 Review effort: Balanced

Comment thread .github/workflows/pr-assess.md Outdated
Port the tested built-in label replacement and standalone-comment behavior.
Keep matching, conflicting, or unreadable outcome labels unchanged.
Limit suggested updates to the PR description, not changes to the code.
Include offline digest-checked probes for the pinned MIT-licensed handler.

Assisted-by: GitHub Copilot (model: GPT-6.1 Sol, autonomous)
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 9ee0ab16-f074-4303-82b9-d11bfad16175
Copilot AI balanced review requested due to automatic review settings October 9, 2026 21:14

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

The new Node-dependent test fails rather than skipping when Node is unavailable.

1 open finding
3 resolved since last review

🧠 Review effort: Balanced

Comment thread tests/test_pr_assess_replace_label.py Outdated
Skip test if Node.js is not available.

Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Stale label state can produce multiple verdict labels, and the Node-dependent test does not follow the repository’s skip convention.

1 open finding
1 resolved since last review

🧠 Review effort: Balanced

Comment thread .github/workflows/pr-assess.md Outdated
Copilot AI balanced review requested due to automatic review settings October 9, 2026 21:26

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

The new label-handler test module has an indentation error that prevents test collection.

2 open findings

🧠 Review effort: Balanced

Comment thread tests/test_pr_assess_replace_label.py Outdated
Follow the extension-submission remove/add pattern: remove up to two stale
outcomes and add the selected outcome only when absent.
Keep matching outcomes unchanged, post fresh standalone comments, and
limit suggested updates to the description.

Remove the obsolete replacement-handler tests and fixtures. Make no
transactional or concurrent-manual-edit guarantee.

Assisted-by: GitHub Copilot (model: GPT-6.1 Sol, autonomous)
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 9ee0ab16-f074-4303-82b9-d11bfad16175
Copilot AI balanced review requested due to automatic review settings October 9, 2026 22:20
@KSchlobohm

Copy link
Copy Markdown
Contributor Author

Updated in 457e1f766c7970a022346fa64dec6b2f88637925 to follow extension-submission's bounded remove/add pattern. Removed the replacement-only tests and fixtures, kept description-only suggestions and fresh comments, and documented the partial-failure/concurrent-edit limits. Local checks: 136 passed, 16 skipped. The race is not claimed resolved.

Prepared on behalf of KSchlobohm with GitHub Copilot (GPT-6.1 Sol, autonomous); AI-assisted change summary.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

The final stability check can publish a verdict using stale PR-title context.

1 open finding
2 resolved since last review

🧠 Review effort: Balanced

Comment thread .github/workflows/pr-assess.md Outdated
Compare title text with the existing captured inputs before reporting.
Require an inconclusive explanation when the title changes during assessment.
Update the existing prompt contract and regenerate its pinned workflow lock.

Assisted-by: GitHub Copilot (model: GPT-6.1 Sol, autonomous)
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 9ee0ab16-f074-4303-82b9-d11bfad16175
Copilot AI balanced review requested due to automatic review settings October 10, 2026 00:04
@KSchlobohm

Copy link
Copy Markdown
Contributor Author

Updated in 0be9efb76e001b86265297ee9f92418025ebc8ba to include the PR title in the final input-stability check and explain title changes as inconclusive. Local checks: 136 passed, 16 skipped; pinned compilation, targeted Ruff, and whitespace passed. Fork live evidence covers a stable-title control and a successful title-change retry.

Prepared on behalf of KSchlobohm with GitHub Copilot (GPT-6.1 Sol, autonomous); AI-assisted change summary.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔵 Needs a closer look

Production behavior depends on LLM orchestration and a large generated workflow, warranting final human review despite strong safeguards and tests.

0 open findings

1 resolved since last review

🧠 Review effort: Balanced

@KSchlobohm
KSchlobohm merged commit 0443760 into github:main Oct 10, 2026
16 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants